Distributed Application Checkpointing for Replicated State Machines
نویسندگان
چکیده
Application checkpointing is a widely used recovery mechanism that consists of saving an application's state periodically to be in case failure. In this study we investigate the utilisation distributed for replicated machines. Conventionally, machines, information stored way each replicas or separately single instance. Applying provides means adjust level fault tolerance approach by giving away from time. We use local cluster and cloud environment examine effects simple machine example compare results with conventional approaches. As expected, gains memory consumption utilise different levels while performing worse terms
منابع مشابه
CASPaxos: Replicated State Machines without logs
CASPaxos is a replicated state machine (RSM) protocol, an extension of Synod. Unlike Raft and Multi-Paxos, it doesn’t use leader election and log replication, thus avoiding associated complexity. Its symmetric peer-to-peer approach achieves optimal commit latency in wide-area networks and doesn’t cause transient unavailability when any bN−1 2 c of N nodes crash. The lightweight nature of CASPax...
متن کاملFast Replicated State Machines Over Partitionable Networks
This paper presents an implementationof replicated state machines in asynchronous distributed environments prone to node failures and network partitions. This implementation has several appealing properties: It guarantees that progress will be made whenever a majority of replicas can communicate with each other; it allows minority partitions to continue providing service for idempotent requests...
متن کاملMencius: Building Efficient Replicated State Machines for WANs
We present a protocol for general state machine replication – a method that provides strong consistency – that has high performance in a wide-area network. In particular, our protocol Mencius has high throughput under high client load and low latency under low client load even under changing wide-area network environment and client load. We develop our protocol as a derivation from the well-kno...
متن کاملApplication controlled checkpointing coordination for fault-tolerant distributed computing systems
In order to provide fault tolerance for distributed systems, the checkpointing technique has widely been used and many researches have been performed to reduce the overhead of check-pointing coordination. In this paper, we present a new checkpointing coordination scheme in which the application controls the coordination activity by utilizing the communication pattern of the application program....
متن کاملCanonical finite state machines for distributed systems
There has been much interest in testing from finite state machines (FSMs) as a result of their suitability for modelling or specifying state-based systems. Where there are multiple ports/interfaces a multi-port FSM is used and in testing a tester is placed at each port. If the testers cannot communicate with one another directly and there is no global clock then we are testing in the distribute...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
ژورنال
عنوان ژورنال: Scalable Computing: Practice and Experience
سال: 2021
ISSN: ['1895-1767']
DOI: https://doi.org/10.12694/scpe.v22i1.1840